Papers with diagnostic tasks
Benchmarking and Mitigating the Impact of Noisy User Prompts in Medical VLMs via Cross-Modal Reflection (2026.eacl-industry)
Copied to clipboard
| Challenge: | Existing medical vision-language models follow user-provided prompts blindly, a new study finds . current models are noisy, causing problems with reliability in real-world interactions . |
| Approach: | They propose a method to evaluate the influence of clinical prompts on medical vision-language models . they use cross-modal reflection chain-of-thought to train the model to produce reasoning paths . |
| Outcome: | The proposed method significantly improves the robustness against noisy prompts . existing Med-VLMs follow user-provided prompts blindly, the authors show . |
FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing approaches to evaluate language models fail to provide structural clarity and verifiable inference. |
| Approach: | They propose to use a large-scale dataset of programmatically verified reasoning traces to evaluate structured logical inference. |
| Outcome: | The proposed model achieves 45.7% accuracy on masked operation prediction and 27% on two-step completion. |
Improving Compositional Generalization with Latent Structure and Data Augmentation (2022.naacl-main)
Copied to clipboard
| Challenge: | Generic unstructured neural networks struggle on out-of-distribution compositional generalization. |
| Approach: | They propose a method to recombinate examples from a model called Compositional Structure Learner and add them to a pre-trained sequence-to-sequence model. |
| Outcome: | The proposed model is even stronger than a T5-CSL ensemble on two real world compositional generalization tasks. |
Good-Enough Compositional Data Augmentation (2020.acl-main)
Copied to clipboard
| Challenge: | a proposed data augmentation protocol provides a compositional inductive bias in conditional and unconditional sequence models. |
| Approach: | They propose a data augmentation protocol that provides a compositional inductive bias in conditional and unconditional sequence models by replacing discontinuous fragments with other fragments that appear in at least one similar environment. |
| Outcome: | The proposed protocol reduces error rate by 87% on diagnostic tasks and 16% on semantic parsing tasks. |
Beyond Static Profiles: Capturing the Fluidity of User Preferences in Diverse Scenarios (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to personalize Large Language Models often default to homogeneous behaviors . preferences can shift, and conflict, depending on context, authors argue . |
| Approach: | They propose a hierarchical taxonomy to differentiate between stable and situational preferences . they use a dataset of 10k meticulously curated preferences to test their taxonomies . |
| Outcome: | The proposed model differentiates between stable and situational preferences based on curated user preferences . it provides a practical testbed for advancing dynamic, context-aware personalization in conversational agents. |
Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational Knowledge (2025.emnlp-main)
Copied to clipboard
| Challenge: | Using large language models, we study their morphosyntactic competence and generalization capabilities. |
| Approach: | They propose to use morphosyntactic tasks to study their linguistic knowledge and generalization capabilities to extract different types of morphological structure for typologically diverse languages. |
| Outcome: | The proposed models outperform GPT-4o and LLaMA 3.3-70B in all diagnostic tasks, but show little evidence of abstract morphological rule learning. |
The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form–Meaning Mapping (2026.acl-long)
Copied to clipboard
| Challenge: | a visual Iconicity test is used to evaluate vision–language models based on visual form and iconicity ratings. |
| Approach: | They propose a video-based benchmark to evaluate vision–language models on three tasks . they assess 17 state-of-the-art VLMs in zero- and few-shot settings on Sign Language of the Netherlands . |
| Outcome: | The proposed benchmark evaluates 17 state-of-the-art VLMs on Sign Language of the Netherlands . they achieve moderate to strong alignment with human iconicity ratings, but fail to infer lexical meaning from visual form alone . |